Papers with German dataset
ADEA: An Argumentative Dialogue Dataset on Ethical Issues Concerning Future A.I. Applications (2024.lrec-main)
Copied to clipboard
| Challenge: | Introducing ADEA: a dataset that captures online dialogues and focuses on ethical issues related to future AI applications. |
| Approach: | They propose a German dataset that captures online dialogues on ethical issues . the dataset includes over 2800 labeled user utterances on four different topics . they use an argument graph as the system's knowledge base and an annotation scheme . |
| Outcome: | The proposed dataset includes over 2800 user utterances on four ethical topics . the aim is to improve knowledge about AI ethics topics through argumentative dialogues . |
Our kind of people? Detecting populist references in political debates (2023.findings-eacl)
Copied to clipboard
| Challenge: | Existing literature on populism has only limited agreement on its exact properties . |
| Approach: | They propose a cross-lingual dataset to identify populist rhetoric in text . they propose 'hierarchical' annotation procedure to annotate populist references . |
| Outcome: | The proposed dataset can be used to investigate how political actors talk about The Elite and The People and to study how populist rhetoric is used as a strategic device. |
GLoHBCD: A Naturalistic German Dataset for Language of Health Behaviour Change on Online Support Forums (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing motivational interviewing methods lack the deep understanding of user utterances that is essential to the spirit of motivational interviews. |
| Approach: | They propose to use a German dataset of naturalistic language around health behaviour change to examine the motivational state of the user. |
| Outcome: | The proposed dataset of naturalistic language around health behaviour change is based on a weight loss forum in germany and is evaluated using theoretically grounded motivational interviewing categories. |
On the Impact of Cross-Domain Data on German Language Models (2023.findings-emnlp)
Copied to clipboard
Amin Dada, Aokun Chen, Cheng Peng, Kaleb Smith, Ahmad Idrissi-Yaghir, Constantin Seibold, Jianning Li, Lars Heiliger, Christoph Friedrich, Daniel Truhn, Jan Egger, Jiang Bian, Jens Kleesiek, Yonghui Wu
| Challenge: | Traditionally, large language models have been trained on general web crawls or domain-specific data. |
| Approach: | They present a German dataset and a dataset aimed at containing high-quality data to examine the importance of data diversity over quality. |
| Outcome: | The proposed model outperforms models trained on quality data on multiple downstream tasks. |
Using Pre-Trained Language Models in an End-to-End Pipeline for Antithesis Detection (2024.lrec-main)
Copied to clipboard
| Challenge: | Rhetorical figures are a "departure from the normal usage" of language . features of metaphors, irony and sarcasm enhance performance of several NLP tasks. |
| Approach: | They propose a pipeline approach to detect rhetorical figures using large language models by splitting text into phrases and identifying parallel phrases with a syntactically parallel structure. |
| Outcome: | The proposed approach outperforms state-of-the-art methods by an F1 score of 65.11 %. |